Back

Frontiers in Genetics

Frontiers Media SA

Preprints posted in the last 30 days, ranked by how well they match Frontiers in Genetics's content profile, based on 230 papers previously published here. The average preprint has a 0.18% match score for this journal, so anything above that is already an above-average fit.

1
Evaluating the Impact of Principal Component and Mixed Model Approaches on Polygenic Risk Score Portability to Diverse Ancestries in the UK Biobank

Harikrishnan, A. S.; Kelly, C. M.

2026-08-19 genetic and genomic medicine 10.64898/2026.08.17.26360388 medRxiv
Top 0.1%
13.0%
Show abstract

Polygenic risk scores (PRS) offer considerable potential for precision medicine. How ever, their predictive performance often attenuates when applied to populations that differ from the genome-wide association study (GWAS) training population. There are many potential sources of this portability problem, and one relatively under-explored contributor is the presence of residual confounding in GWAS summary statistics. In particular, confounding specific to the training population may contribute to predictive performance that does not transfer to other populations, such that improved control of population stratification could potentially improve PRS portability. Here, we investigated whether varying levels of population stratification adjustment, through the inclusion of principal components and the use of mixed models, altered PRS portability in three broad ancestry groups in the UK Biobank. The PRS were built using European training data for coronary artery disease and type 2 diabetes and subsequently evaluated in South Asian, African, and Latin American participants. We found that increasing PC adjustment did not produce a consistent trend in portability across ancestry groups or phenotypes, despite modest reductions in the LDSC intercept. However, substantial ancestry- and phenotype-specific effects on transferability were observed. Mixed-model association provided no significant change in PRS discrimination or portability. These findings highlight the need for a better understanding of the nature of residual confounding in PRS and whether improving the causal validity of GWAS results can ultimately improve the transferability of predictive accuracy between populations.

2
Uncovering High-Order Epistatic Interactions in GWAS via a Machine Learning-Based Feature Engineering Framework

Byun, J.; Saha, D.; Han, Y.; Shaw, V. R.; Siminovitch, K.; Amos, C. I.

2026-08-09 genomics 10.64898/2026.08.03.742638 medRxiv
Top 0.1%
11.8%
Show abstract

BackgroundGenome-wide association studies (GWAS) often fail to identify higher-order epistatic interactions that contribute to complex inheritance patterns of traits and diseases. While machine learning (ML) can capture non-linear relationships, extracting interpretable insights from these models remains a challenge. We propose a novel tree-based feature engineering framework that uses Classification and Regression Trees (CART) to explicitly encode high-order interaction decision paths as dummy variables. We investigate three path-based encoding strategies: (i) all decision paths, (ii) leaf-node paths only, and (iii) internal-node paths only. This approach aims to transform complex decision boundaries into discrete features that capture nonlinear interactions that are not readily captured by traditional association models. ResultsThe framework was evaluated using genetic data for ANCA-associated vasculitis (AAV). To manage the high dimensionality of the engineered feature space, we applied a comprehensive suite of ML methods across three tasks: (1) Ensemble Learning (Random Forest, XGBoost, and Gradient Boosting Machine); (2) Decision Tree Analysis (CART); and (3) Regression and Classification Tasks (Regularized Linear Regression/LASSO, Support Vector Machine, and Logistic Regression). Stepwise feature selection and regularization were employed to isolate the most informative interaction patterns. Results indicate that incorporating CART-derived interaction paths--particularly those from high-impact regions of the tree--significantly improves classification accuracy and model interpretability compared to using the original feature space alone. ConclusionsThe proposed framework provides a robust, scalable methodology for identifying high-order genetic interactions. By bridging the gap between the predictive power of ensemble ML and the necessity for mechanistic insight, this approach offers a clearer mapping of the combinatorial genetic processes underlying complex diseases. While applied here to AAV, the method is highly adaptable for exploring the genetic architecture of diverse populations and complex traits.

3
Genome-Wide Selection Signatures in Nili-Ravi Buffalo (Bubalus bubalis) Reveal a T-Cell Costimulatory and Cytokine-Signaling Gene Network Distinct from Classical Bovine Tuberculosis Candidate Genes

Ahmad, A.; bakar, A.; Laeeque, S. M.; Khan, W. A.; Kaul, H.; Manan, A.; mustafa, h.

2026-08-11 genomics 10.64898/2026.08.10.743898 medRxiv
Top 0.2%
8.0%
Show abstract

Genomic signatures of selection can reveal loci underlying adaptation and disease resistance in livestock populations, but such analyses in water buffalo (Bubalus bubalis) have historically been constrained by the absence of a chromosome-level, species-native reference genome for SNP array data. We re-analyzed genotype data from 85 Nili-Ravi buffalo (Axiom Buffalo Genotyping 90K array, originally positioned using bovine (Bos taurus, UMD3.1) proxy coordinates, by performing a full coordinate liftover to the buffalo-native UOA_WB_1 assembly using an independently published SNP remapping resource. Following quality control (51,209 markers retained), haplotype phasing, and genome-wide integrated haplotype score (iHS) and Wrights Fst (case/control) selection scans, we evaluated 14 classical bovine-tuberculosis (bTB) candidate genes and identified six additional genes with putative immune function through an unbiased genome-wide screen. None of the 14 classical candidates (including SLC11A1, the Toll-like receptors, and IFNG) reached genome-wide significance in either scan. In contrast, six novel loci TNFSF18, IL2RB, TNFRSF19, IRF2, IL15, and CD28 showed significant iHS or Fst signals, four of which (TNFSF18, IL2RB, IL15, CD28) converge functionally on T-cell costimulation and cytokine receptor signaling (KEGG pathways map04660 and map04060, Bos taurus proxy annotation). Using extended haplotype homozygosity (EHH) decay, haplotype furcation structure, and per-marker haplotype counts as three independent lines of corroborating evidence, we classified these six genes into confidence tiers: TNFSF18 and IL2RB showed the strongest, most balanced support, while CD28 and IL15 signals were driven by very few haplotypes (3 and 5 of 30, respectively) and should be interpreted cautiously pending replication. These findings suggest that adaptive, cell-mediated immune signaling rather than the innate/macrophage-centred mechanisms emphasized by existing bTB candidate gene panels may be a more productive avenue for future selection studies in Nili-Ravi buffalo, while underscoring the value of buffalo-native coordinate systems for accurate genomic inference in this species.

4
Genetic Architecture and Sample Size Impact Relative Performance of Nonlinear Machine Learning and Standard Polygenic Risk Scores

Zhu, J.; Baousi, A.; Morris, A. P.; Guo, H.

2026-09-03 genetic and genomic medicine 10.64898/2026.08.29.26361109 medRxiv
Top 0.3%
6.6%
Show abstract

Standard polygenic risk scores (PRSs) are constructed based on additive genome-wide association study (GWAS) summary statistics. Nonlinear machine learning methods have been increasingly applied to construct PRSs directly from individual-level data, with the aim of improving predictive performance over standard PRSs through their ability to model non-additive genetic effects. However, their superiority across studies has been inconsistent, and the conditions under which they provide meaningful improvements remain unclear. We combined theoretical analysis, simulations and a real-world application to investigate when two widely used nonlinear machine learning methods, random forest and XGBoost, outperform standard PRSs. Theoretical analysis showed that standard PRSs can implicitly capture part of the genetic variance attributable to nonadditive genetic effects through their contributions to marginal SNP effects, thereby losing less information than commonly assumed. Although nonlinear models have a higher theoretical potential, their greater flexibility incurs a bias-variance trade-off that can limit predictive gains at finite sample sizes. Simulations showed that XGBoost outperformed the standard PRS only when the genetic architecture involves a sufficiently large proportion of interaction genetic variance concentrated across relatively few interaction effects and large training samples were available. Random forest consistently underperformed the standard PRS. In an application to ischemic heart disease prediction using UK Biobank data, XGBoost showed no meaningful improvement in predictive performance over the standard PRS, whereas random forest again performed worse. Together, these findings suggest that nonlinear machine learning do not uniformly outperform standard PRSs; rather, their relative performance depends jointly on genetic architecture and training sample size. Our study helps to reconcile the inconsistent results reported across previous studies and provides a framework for identifying settings in which more complex PRS models are likely to be beneficial.

5
Benchmarking Imputation Methods for Single-Cell RNA Sequencing Data Using Peripheral Blood Mononuclear Cells from Acute Myocardial Infarction Patients

Ramesh, P.; Fyta, M.

2026-08-27 bioinformatics 10.64898/2026.08.23.746230 medRxiv
Top 0.3%
6.5%
Show abstract

Acute myocardial infarction (AMI) remains one of the leading causes of mortality worldwide, and the following post-effects, such as post-AMI inflammation and tissue repair, involve peripheral blood mononuclear cells playing a critical role. The influence of imputation methods in biological data is assessed with respect to high-resolution single-cell RNA sequencing (scRNAseq) data relevant to these cells. Still scRNAseq data often encounter a lot of dropout events, leading to sparse and noisy datasets, hampering downstream results. To assess the influence of the missingness in the data, we artificially impose different levels of dropout in available scRNAseq data by leveraging various imputation techniques. Specifically, we introduce artificial missingness at 10%, 20%, and 30% levels under a missing completely at random (MCAR) framework, repeated across 10 independent runs. We benchmarked six imputation strategies - MAGIC, IterativeImputer, KNNImputer, Mean Imputation, SoftImpute, and a Generative adversarial network (GAN) - based approaches using multiple evaluation metrics: marker gene preservation, clustering consistency (Adjusted Rand Index - ARI), gene-wise correlation with ground truth, and structural separation (silhouette scores). The results clearly underline that no single imputation method dominated across all metrics. Overall, Mean and KNN imputers showed limited recovery across all benchmarks. GAN excelled in global transcriptional recovery and SoftImpute in preserving biologically meaningful cell-type signals. Our results highlight the importance of selecting the imputation methods as part of the pre-processing step towards the downstream biological questions related to transcriptome recovery, detection of marker genes, or maintaining cell-type-specific resolution.

6
Prioritizing Genes and Rare Protein-Coding Variants in Acute Myeloid Leukemia via Whole Genome Sequencing Data

Vieno, S.; Singh, M.; Kramer, S.; Chatzinakos, C.; Peterson, R.; Riley, B.; Bacanu, S.-A.; Dinh, T.; Trinh, B. Q.; Nguyen, T.-H.

2026-08-22 genetic and genomic medicine 10.64898/2026.08.19.26360760 medRxiv
Top 0.4%
5.4%
Show abstract

The extent to which rare and common genetic variants jointly contribute to the risk of acute myeloid leukemia (AML) still remains relatively unexplored in large-scale biobank whole-genome sequencing cohorts. Here, we leverage the latest sequencing and phenotypic data from the All of Us Research Program to identify variants, genes, and gene-sets associated with AML. We performed set-based association tests for rare protein-coding variants (Ncases=265 and Ncontrols=169,706) and single-variant association tests for common variants (Ncases=265 and Ncontrols=169,705) utilizing the large European-like ancestry sample. For the rare-variant set-based tests conducted using SAIGE-GENE+, four genes were statistically significant: DNMT3A, TET2, SRSF2, and IDH2 (Bonferroni-corrected Cauchy p-value < 0.05). We also constructed multiple rare-variant burden risk scores using different gene-sets to identify those with a substantial rare-variant burden for AML. Gene-sets derived from Genomic Data Commons whole-genome sequencing data, comprising two distinct groups-genes observed to harbor somatic mutations in AML and genes observed to harbor somatic mutations across all cancer types-showed a statistically significant rare-variant burden (Bonferroni-corrected p-value < 0.05). Ultimately, these findings demonstrate that leveraging whole-genome sequencing in large-scale biobanks enables the identification of rare protein-coding variants, genes, and gene sets associated with AML.

7
A Practical Framework for Constructing Population-Specific and Alternate-Contig-Aware Genome References: A case study of Vietnam

Vo, N. S.; Tran, T. T. H.; Duong, V. C.; Nguyen, N. N.; Pham, T. M.; Vu, Q. T.; Tran, M. H.; Hoang, T. H.; Nguyen, Q.; Nguyen, D. T.

2026-08-27 genomics 10.64898/2026.08.24.746817 medRxiv
Top 0.5%
5.0%
Show abstract

Current studies in human genomics typically rely on the standard genome reference GRCh38 which is known to be biased toward populations of European ancestry and therefore has limitations when applied to other populations. Although various graph-based pangenome references were constructed for several populations to deal with this bias, their usage in practice is currently still limited compared to linear genome references. Here we present a framework for constructing a population-specific genome reference using GRCh38 as backbone with alternate-contig awareness to enhance genomic data analysis in the target population. We demonstrated the advantages of our framework using both public and in-house Vietnamese whole-genome sequencing (WGS) datasets. Genomic variants derived from high-coverage WGS data of the 1000 Vietnamese Genomes Project (VN1K) were imported into our framework to build a Vietnamese-specific Genome Reference (VGR). VGR was then compared to GRCh38 in read alignment and variant calling using high-coverage WGS data of 99 Vietnamese individuals (KHV) from the 1000 Genomes Project (1kGP). Using Omni array genotyping data from 99 KHV samples as an independent benchmark, we found that VGR improved variant-calling precision and reduced false-positive calls compared to GRCh38. Our framework could be easily used for other populations as long as they have a variant database similar to VN1K. Our code is publicly available at github.com/VinGenome/VGR

8
African Green Monkey Cerebrospinal Fluid miRNome Captures Conserved miRNAs Relevant to Human Neurodegenerative Disease

Dzigurski, S.; Al-Abri, R.; Li, X.; Grasty, M. R.; Rodrigues, A. C.; Weed, M. R.; Elsworth, J. D.; Lawrence, M. S.; Heng, Y. J.; Bogsan, C. S.; Naderi Yeganeh, P.; Hide, W. A.; Slack, F. J.; Gursoy, G.; Miranker, A. D.; Brown, B. R. P.

2026-08-13 genomics 10.64898/2026.08.07.743104 medRxiv
Top 0.5%
5.0%
Show abstract

BackgroundThe African green monkey (AGM) is increasingly used as a model for early-stage Alzheimers disease (AD), with cerebrospinal fluid (CSF) targeted for biomarker discovery and longitudinal disease monitoring of shifts in the central nervous system. MicroRNAs (miRNAs) are particularly informative indicators of early neuropathological change. Despite the complementary value of an early-stage disease model and a molecular marker capable of capturing early change, the miRNA composition (miRNome) of AGM remains undefined. We established the AGM CSF miRNome from antemortem samples using miRNA sequencing and a qRT-PCR-based array. We also developed a hierarchical annotation pipeline to classify miRNAs as either family-conserved or unclassified and to assess sequence alignment across humans and other species. ResultsWe used untargeted miRNA sequencing to characterize the AGM CSF miRNome and identified 205 miRNAs that could be classified into three family-conserved categories: canonical, noncanonical, and 3'-terminal variants. Of these, 150 were also detected using a human-targeted qRT-PCR array, providing independent support for the sequence-derived miRNome. Sequencing abundance and qRT-PCR array Ct values showed significant cross-platform concordance overall, although concordance was lower for 3'-terminal isomiRs than for canonical miRNAs. Comparison with human GTEx tissue-expression data indicated that several human homologs of AGM CSF miRNAs exhibited brain-preferential expression. Notably, predicted targets of many of these miRNAs were enriched for pathways implicated in neurodegenerative disease. Finally, we identified 20 unclassified candidates that could not be assigned to established miRNA families, two of which we propose as putatively novel miRNAs. ConclusionThe AGM CSF miRNome is substantially conserved with the human miRNome but also contains 3'-terminal isomiRs and unclassified miRNA candidates. AGM CSF contains miRNAs homologous to human miRNAs associated with AD and other neuropathologies, highlighting the translational potential of this model. However, our study also reveals challenges related to species-specific sequence variation and reduced cross-platform concordance for isomiRs. Thus, comparative studies will be needed to validate the functional and biomarker relevance of these miRNAs across species. More generally, this initial miRNome provides a reference resource for future studies of miRNAs in AGM across disease-related, physiological, experimental, and evolutionary contexts.

9
Genome sequencing reveals novel pathogenic deep-intronic PCDH15 variants, amenable to antisense oligonucleotide-based splice correction

Rodenburg, K.; Fenwick, L.; Pennings, R.; Haer-Wigman, L.; Ben-Yosef, T.; van Erp, F.; Reurink, J.; Gilissen, C.; van den Born, L. I.; Cremers, F. P. M.; Cohen, Y.; Yntema, H.; de Vrieze, E.; Kremer, H.; de Bruijn, S. E.; Collin, R. W. J.; Roosing, S.; van Wijk, E.

2026-08-24 genetics 10.64898/2026.08.20.746067 medRxiv
Top 0.6%
4.8%
Show abstract

Despite substantial advances in diagnostic testing, 10-15% of Usher syndrome patients remain without a genetic diagnosis, having significant implications for genetic counseling and potential future therapeutic interventions. In this study, genome sequencing data from probands clinically presenting with Usher syndrome were analyzed. Two novel deep-intronic variants were identified in PCDH15, c.3983+3635A>G and c.3123-1728A>G, in two independent patients. Both deep-intronic variants were classified as likely pathogenic and predicted to alter PCDH15 pre-mRNA splicing. Using a minigene splice assay and iPSC-derived photoreceptor precursor cells from patients, we confirmed that both variants lead to the inclusion of a pseudoexon in the PCDH15 transcript introducing a stop codon and subsequent premature termination of protein translation. We designed and evaluated antisense oligonucleotides (ASOs) with the purpose of redirecting aberrant pre-mRNA splicing caused by both deep-intronic variants. For both variants, designed ASOs were successful in restoring normal splicing patterns, highlighting their potential as a future therapeutic intervention strategy to halt the progression of retinitis pigmentosa caused by these novel variants. Overall, these findings contribute to the understanding of Usher syndrome caused by deep-intronic pathogenic variants in PCDH15 and describe for the first time the use of an ASO-mediated splice correction strategy for individuals diagnosed with these variants.

10
"Transcriptional and isoform-level regulation of lipid-candidate genes in preeclamptic placentas"

Eyer, K. S.; Lemaire, M.; Fan, X.; Wilson, S. L.

2026-08-21 genomics 10.64898/2026.08.17.745256 medRxiv
Top 0.6%
4.8%
Show abstract

Preeclampsia (PE) is a hypertensive pregnancy-specific disorder and a leading cause of maternal and fetal mortality. A common feature of PE placentas and maternal plasma is dyslipidemia, or abnormal lipid levels, which can increase oxidative stress and endothelial dysfunction. However, the precise transcriptional, post-transcriptional, and epigenetic mechanisms underlying these abnormalities remain poorly characterized. Identifying such changes may clarify disease mechanisms and identify lipid-related PE biomarkers. We conducted a large-scale meta-analysis integrating public placental datasets from NCBI GEO, comprising four DNA methylation (DNAm) datasets (n = 172), three RNA-sequencing datasets (n = 92), and an independent RNA microarray validation cohort (n =146). We evaluated differential DNAm (limma), gene expression (DESeq2), transcript-level shifts (Swish), and alternative splicing (rMATS) in PE versus control placentas, with all analyses stratified by fetal sex via an interaction term model. We also performed placental cell-type deconvolution to quantify PE-associated cell-type proportion changes. Our results demonstrated that lipid-related regulation changes in PE placentas occur primarily at the gene and transcript level, with DNAm showing no changes. We also identified significant isoform switching in PE that were undetected by differential gene expression analysis, and primarily driven by alternative transcription initiation and termination sites rather than alternative splicing. A subset of these isoform switches mapped to pathways dysregulated in PE and were predicted to cause functional protein changes. An interaction term model identified several sex-specific differentially expressed genes (DEGs) in PE, including a subset of male-specific downregulated genes involved in oxidative metabolism. However, many of the remaining sex-specific DEGs across both sexes were previously uncharacterized in the literature. These findings suggest that transcriptional and isoform-level regulation play a role in PE-associated dyslipidemia, with certain regulatory pathways displaying fetal sex-specific patterns. Highlights- Preeclampsia-associated dyslipidemia manifests at the gene and transcript level - Reciprocal isoform switches were missed by standard gene-level analyses - Alternative transcript initiation and termination drove isoform switching - Sex-interaction modeling identified sex-specific transcriptional shifts in PE

11
Chromosome assembly for the Black bean aphid Aphis fabae

Whitehead, M. A.; Claudia Wierzbicki, C.; Hughes, M.; Darby, A. C.

2026-08-11 genomics 10.64898/2026.08.05.743085 medRxiv
Top 0.6%
4.8%
Show abstract

The black bean aphid, Aphis fabae is a crop pest and vector of insect-transmitted pathogens, comprising closely related sub-species with overlapping host ranges. In other Aphis species, over-expression of specific detoxification genes has been linked to insecticide tolerance. We present two chromosome-scale assemblies for a clonal A. fabae line, representing two phased haplotypes, generated using HiFi and Hi-C sequencing technologies. A comprehensive genome annotation, built with PacBio Iso-Seq data, was used to investigate genes underlying insecticide tolerance. Both genomes are comprised of four chromosomal blocks (haplotype 1: 427 Mb; haplotype 2: 396 Mb) with high BUSCO completeness (98.7%). Comparative genomics revealed an expansion of UDP-glycosyltransferases, whose expression is linked to insecticide detoxification in other Aphis species. These high-quality references provide a foundation for studying A. fabae sub-species and a genomic resource for investigating insecticide tolerance across the Aphis genus. Author summaryHere we have provided a comprehensive assembly and annotation for further study into the Black bean aphid, Aphis fabae, using up to date long-range sequencing technologies. The final assemblies for both haplotypes are chromosome length and consist of 4 main chromosome blocks, consistent with the literature. The A. fabae genome was found to contain an increase in copy number of UDP-glycosyltransferases, which have previously been linked to insecticide resistance. The work here will be a resource to those studying insecticide tolerance in crop pests, as well as the differences between A. fabae sub-species.

12
Uveal and cutaneous melanoma share a common mutation with distinct prognostic implications: A bioinformatic study

Razmjooei, F.; Ashayeri, H.; Jafarzadeh, Z.; Dabbaghabdollahi, P.; Jafarizadeh, A.

2026-08-11 genetic and genomic medicine 10.64898/2026.08.07.26359988 medRxiv
Top 0.6%
4.8%
Show abstract

Background: Uveal melanoma (UM) and cutaneous melanoma (CM) both originate from the same cell line. This proposes the possibility of a shared mechanism between entities, requiring explicit investigation. Methods: Data from GWAS Catalog and DisGeNET were used to identify shared variation-disease associations (VDAs) between UM and CM. The results were validated using the Ensembl database. In the next step, the STRING database was used to identify the protein-protein interaction. Results: Subsequently, 109 unique VDAs were identified for UM and 880 for CM. However, only 2 VDAs were found to be shared among UM and CM in different ethnic groups. These shared VDAs were rs12203592 of the IRF4 gene, rs12913832 of the HECT and RLD domain-containing E3 ubiquitin protein ligase 2 (HERC2) gene. Notably, PPI network assessment through STRING showcased that OCA2 and IRF4 directly interacted with HERC2. Conclusion: While HERC2 acts as a poor prognostic factor in uveal melanoma, IRF4 status is a key prognostic indicator in both UM and CM. Identifying IRF4 allele contributions enables a better understanding of melanoma pathogenesis and fosters the development of disease-specific approaches.

13
Whole-genome resequencing-based comparative variant analysis identifies candidate genes associated with cross-beak phenotype in Huiyang Bearded chickens

Ye, F.; Yu, H.; Hong, Y.; Zhao, H.; Kang, H.; Yu, H.; Li, H.

2026-08-18 genomics 10.64898/2026.08.11.744104 medRxiv
Top 0.6%
4.5%
Show abstract

Cross-beaks are deemed a threat to poultry health, productivity, and animal welfare. Nevertheless, due to sporadic cases, heterogeneity of gene loci and incomplete dominance, the molecular mechanism of cross-beak formation, especially the degree of cross, is not yet clear. Thus, we screen key genes and reveal the possible phenotypic formation mechanism of cross-beak by comparison with different degrees of deformity in Huiyang Bearded chickens by compare whole-genome resequencing-based variant analysis. Comparative analysis between cross-beak and normal-beaked chickens identified differential variants in several candidate genes, including CDH11, CTNNAL1, NRXN3, NRXN1, CDH5, SDC3, and DHFR. Genes harboring these variants were enriched in pathways related to cell adhesion molecules and metabolic processes, with functional annotations involving cell-cell adhesion and neural crest cell migration. Comparative analysis between chickens with severe and slight cross-beak deformities identified additional candidate genes, including MRPL21, NSUN2, DDX55, GNB3, and NFKB2. These genes were associated with enriched terms and pathways related to focal adhesion, amyotrophic lateral sclerosis, steroid 7 -hydroxylase activity, and skin-barrier establishment. These findings provide a preliminary catalogue of genetic variants and candidate genes for future functional studies of cross-beak development and severity in chickens.

14
Benchmarking Graph Neural Networks for Multi-Omics Cancer Subtyping using Methylation and Gene Expression Profiles

Schirmacher, J.; Maurer, M. C.; Metsch, J. M.; Ploesch, S.; Chereda, H.; Blumenthal, D. B.; Hauschild, A.-C.

2026-08-25 bioinformatics 10.64898/2026.08.21.745839 medRxiv
Top 0.8%
4.1%
Show abstract

Motivation: Graph Neural Networks (GNNs) have gained increasing interest in the biomedical domain, as the integration of prior knowledge and deep neural networks has the potential to enhance insights into molecular processes and disease mechanisms. However, a comprehensive and systematic assessment of model architectures, data modalities, graph structures, and their performance for graph signal classification in the biomedical domain is yet to be performed. In order to close this gap, we conducted a benchmarking study on multiple GNNs on a Protein-Protein Interaction (PPI) network for Kidney Renal Clear Cell Carcinoma and Breast cancer subtype prediction, performing an in-depth investigation of architectures, incorporating skip connections and various data modalities. Results: While none of the GNNs outperforms the structure-agnostic Multi-Layer Perceptron baseline, all of them can handle bimodal data (gene methylation and expression) and offer the ability to gain explainability based on PPIs. We offer practical guidelines for applying GNNs to graph signal processing tasks specifically for cancer classification. Depending on the underlying dataset and PPI structure employed, models on different data modalities outperform others. Overall, we suggest using ChebNet, which tends to outperform the Graph Convolutional Network and the Graph Attention Network in cancer subtype prediction. We recommend using GNN architectures that employ a simple flattening readout layer, as they provide better classification performance and faster training time than those with global average pooling. Additionally, we tested residual connections, but they had only an insignificant impact on classification performance.

15
A meta-analysis of ancient and present-day Central Eurasian genome data to revise archaic hominin ancestry

Rymbekova, A.; Kuhlwilm, M.

2026-08-14 genomics 10.64898/2026.08.10.743976 medRxiv
Top 0.9%
3.9%
Show abstract

Archaic introgression has shaped the evolutionary history of Eurasian populations, yet Central Eurasian region remains understudied despite being at the crossroads of ancient human migration. Here, we analyzed the whole-genome data of five Central Eurasian (CE) individuals from Early Bronze Age (EBA) and five present-day CE individuals to characterize the archaic introgression landscape. We estimated that archaic introgression from Neanderthal and Denisovan archaic hominins comprises approximately 2.2% of the Central Eurasian genomes. Both amount and chromosomal distribution of archaic introgression remained largely unchanged between the EBA and present-day CE individuals. Putative introgressed fragments matching the Altai Neanderthal and the Altai Denisovan were retrieved. Our results suggest that while the archaic introgression levels seemingly remained stable over the past several thousand years, larger modern CE genomes panels will be required to fully characterize the genomic landscape of archaic ancestry in the region.

16
Bridging Morphology and Genomics: A rapid image-based assessment of genomic admixture in the endangered gayal (Bos frontalis)

Ma, J.; Chen, Y.; Guo, Z.; Xiao, J.; Wu, H.; Luo, J.; Zhang, Y.-p.; Li, Y.

2026-08-25 zoology 10.64898/2026.08.25.746947 medRxiv
Top 0.9%
3.7%
Show abstract

Abstract The gayal (Bos frontalis) is an endangered semi-domesticated bovine species renowned for its high-quality beef. However, its semi-feral lifestyle, ongoing habitat fragmentation, and extensive genetic introgression from sympatric local cattle have led to dramatic population decline and severe erosion of purebred genetic integrity, posing substantial challenges to its conservation and utilization. To address the urgent demand for rapid, non-invasive, and field-compatible germplasm identification, we developed an integrated artificial intelligence (AI) framework that predicts genomic admixture composition from external morphological images. We constructed a comprehensive dataset comprising 6,245 morphological images and matched genomic sequences from 52 gayals maintained at the Yunnan Provincial Gayal Conservation Farms. Following a preliminary evaluation of nine deep learning models, five were incorporated into a anatomical segment-based multi-modal pipeline, among which Inception_V3 delivered the optimal overall performance. To enhance simultaneous extraction of local fine-grained features and global structural information, we further designed an innovative HybridInceptionViT model by integrating the multi-scale Inception module with the Vision Transformer (ViT) framework. This hybrid model significantly outperformed the baseline Inception_V3, boosting the accuracy of phenotype-derived prediction against genomic admixture estimate from 69.69% to 87.87% (absolute error <15%). This study establishes a practical, low-cost "phenotype-to-genotype" tool for rapid on-site gayal germplasm screening, offering a scalable strategy for the conservation and breeding management of endangered livestock, and holds broad application prospects for agricultural and livestock production systems.

17
Comprehensive analysis of mulberry genetic diversity based on 1-DNJ content and SNP markers

Shen, Z.; Li, J.; Shi, J.; Li, Z.; Wang, F.; Geng, J.; Hu, K.

2026-08-19 genetics 10.64898/2026.08.11.744330 medRxiv
Top 0.9%
3.6%
Show abstract

Mulberry trees have high economic and ecological value, and a robust molecular marker system plus germplasm genetic diversity analysis is critical for innovative utilization of high-quality medicinal and economic mulberry germplasm. Here, 51 mulberry samples were used to develop SNP primers via genome resequencing, with the SNP-PCR system optimized by single-factor and orthogonal assays. The phenotypic diversity and SNP molecular marker genetic diversity of 1-deoxynojirimycin (1-DNJ) in mulberry leaves were analyzed respectively, and the genetic correlation between molecular markers and phenotypic traits was evaluated by Mantel test. Tested germplasm showed marked 1-DNJ variation (0.4805-2.5300 mg/g, CV=0.4241), reflecting rich genetic diversity. The optimal SNP-PCR system included Buffer (containing Mg{superscript 2}+) 2.2 L, 2.5 mM dNTP 0.4 L, forward and reverse primers (10 mol{middle dot}L-1) totaling 2.75 L, Taq DNA polymerase (5 U{middle dot}L-1) 0.3 L, DNA (50 ng{middle dot}L-1) 1.1 L, and ddH2O 13.65 L. 23 highly polymorphic ones amplified 91 loci (81 polymorphic, 89.10% polymorphism rate). Genetic diversity analysis showed that the average genetic distance was 0.3010, and the average expected heterozygosity (H) and Shannon information index (I) reached 0.4667 and 0.3104 respectively, indicating that the genetic differentiation among the tested mulberry germplasms was significant and the population had a moderate to upper level of genetic diversity. UPGMA clustering divided 51 germplasms into 6 major groups at a genetic similarity coefficient of about 0.7, while phenotypic clustering based on 1-DNJ content divided them into 2 major categories and 4 subcategories, with high 1-DNJ germplasm clustered independently. Mantel correlation analysis showed that 6 SNP sites were significantly weakly correlated with 1-DNJ content (r < 0.3, p < 0.05), and can be used as candidate molecular markers for subsequent genetic analysis of 1-DNJ content.This study established a stable mulberry SNP-PCR system, Analyze the molecular genetic characteristics of mulberry germplasm and DNJ phenotypic variation rules respectively, and provide basic data for cluster comparison. and provided a scientific basis for marker database improvement, germplasm identification and molecular-assisted breeding.

18
Detecting CYP2C19 deletions from genotyping array signals using neural networks

Yelmen, B.; Hofmeister, R. J.; Lutsar, V. K.; Finianos, M.; Stone, B. C.; Joeloo, M.; Krebs, K.; Kivistik, P. A.; Smit, S.; Estonian Biobank Research Team, ; Metspalu, M.; Hudjashov, G.; Milani, L.

2026-08-25 bioinformatics 10.64898/2026.08.21.746170 medRxiv
Top 0.9%
3.5%
Show abstract

Since copy number variations (CNVs) in pharmacogenes can cause significant alterations in drug metabolism, their reliable detection is of high importance both for large-scale studies and personalized medicine. Whole-genome sequencing, and specifically long-read sequencing, is the gold standard for CNV detection. Despite increasing availability of these technologies, genotyping arrays are still widely used as cost-effective alternatives in biobank and clinical settings, yet calling CNVs based on array intensity signals is challenging due to low base pair resolution. In this work, we developed a neural network model, nnCNV, to predict deletions in the CYP2C19 pharmacogene region from array intensity signals. We compared our method to the most widely used algorithm, PennCNV, and demonstrated better performance reaching 100% accuracy in the test dataset. Furthermore, we predicted probe-by-probe CYP2C19 deletion coordinates for all Estonian Biobank samples using nnCNV and PennCNV, and validated these predictions using an identity-by-descent (IBD) sharing method, which also demonstrated superior nnCNV performance. For the deletion samples with conflicting PennCNV and nnCNV predictions, we performed PCR analysis for validation, which showed 97% precision for nnCNV compared to 23% for PennCNV. Finally, we assessed the gradient-based feature importance maps and showed that nnCNV utilizes signal intensity information not only from deletion probes, but also from probes in flanking regions. Our results demonstrate that long-range information, which cannot be utilized by hidden Markov models, can improve CNV calling.

19
When the Background Matters: Topic-Dependent reference lists in GWAS and Exome Analyses

Timoney, B.; Guasoni, P.; Zade, K.; Bach, S.; Tropea, D.

2026-08-21 bioinformatics 10.64898/2026.08.14.744838 medRxiv
Top 0.9%
3.5%
Show abstract

Gene Ontology (GO) Biological Process overrepresentation analysis is widely used to interpret gene lists from genetic studies, yet results depend critically on the background (universe/reference list) against which enrichment is tested. This paper examines how genome-exome background mismatch alters GO Biological Process significance and induces annotation-driven bias. First, Monte Carlo simulations across multiple input gene list sizes show that enrichment p-values shift systematically when lists sampled from an exome-like universe are tested against a genome background (and vice versa), producing both inflation and deflation of significance depending on GO term composition; these shifts increase with gene list size. Second, applied analyses of gene lists derived from Genome-Wide Association Studies (GWAS) and Whole Exome Studies (WES) across brain, immune, and metabolic domains demonstrate that background choice changes the set of significant GO IDs, yielding reference-specific terms consistent with both Type I errors (false positives) and Type II errors (false negatives). Because genome backgrounds are commonly used by default, the practical risk is greatest when WES-derived lists are analyzed with genome reference lists. To support reproducible best practice, we provide a simple command set for selecting and documenting study-appropriate backgrounds and for assessing sensitivity of GO Biological Process results to the chosen universe.

20
Evaluating Aggregated Gene Level eQTL Scores

Meyer, D.; Popko, N.; Laub, D.; Schofield, P.; Amariuta, T.; Alexandrov, L. B.; Carter, H.

2026-08-26 bioinformatics 10.64898/2026.08.21.746287 medRxiv
Top 0.9%
3.4%
Show abstract

Genetic feature engineering, used in methods such as transcriptome-wide association study, supports gene-trait association testing by aggregating single variants into gene-level features predictive of expression. To evaluate how different model architectures, LD filtering thresholds, and variant prioritization methods affect expression prediction quality, we trained over 3 million models and evaluated their performance in independent cohorts. Using the best performing models to impute expression and immunotherapy response as an example trait, we found a significant association with the reactive oxygen species pathway (p=0.032). Our model training workflow will support genetic feature engineering towards improved complex trait modeling.